Back

Journal of Chemical Information and Modeling

American Chemical Society (ACS)

Preprints posted in the last 30 days, ranked by how well they match Journal of Chemical Information and Modeling's content profile, based on 238 papers previously published here. The average preprint has a 0.19% match score for this journal, so anything above that is already an above-average fit.

1
Benchmarking Docking Protocols for GPCR Allosteric Modulators

Thompson, T. D.; Miao, Y.

2026-08-20 bioinformatics 10.64898/2026.08.12.744492 medRxiv
Top 0.1%
71.0%
Show abstract

G protein-coupled receptor (GPCR) allosteric modulators (AMs) offer significant therapeutic advantages over orthosteric drugs, yet structure-based virtual screening lacks validated protocols accounting for the conformational complexity of GPCR allosteric sites. We benchmark docking protocols using PDB experimental structures and structural ensembles derived from Gaussian accelerated Molecular Dynamics (GaMD) simulations across four Class A GPCRs (including the muscarinic M2 and M4 receptors, the {beta}2-adrenergic receptor, and the C-C chemokine receptor type 2) with four programs (Glide HTVS, AutoDock Vina, DOCK3.8, and Boltz-2) against experimentally validated modulator libraries and property-matched decoys. GaMD ensemble docking improved early AM enrichment across all four targets under at least one program. Glide ensemble docking was the only protocol to consistently improve early AM recovery across all four targets, ranking known actives almost exclusively within the top 0.5% of compounds at CCR2 and improving M2R active recovery nearly 9-fold relative to the PDB structure. GaMD free-energy landscape topology governed ensemble re-ranking strategy selection: population-skewed landscapes favored top binding energy ranking (BEmin) while flat, multi-populated landscapes favored average binding energy ranking (BEavg), and at targets with dominant low-energy states, a single GaMD cluster matched or exceeded full ensemble or PDB performance. Taking the union of top percentile hits identified by both ensemble re-ranking methods, BEmin / BEavg, maximizes chemical diversity at the earliest percentiles. Program-specific scaffold recovery biases further motivated a consensus BEmin / BEavg approach to maximize hit diversity. The Boltz-2 deep-learning program showed minimal sensitivity to GaMD templates and underperformed conventional docking, suggesting its affinity predictions complement rather than replace physics- and empirical-based docking approaches for GPCR AM screening.

2
Structural Context Determines Docking Engine Performance: A Family-Stratified Benchmark of Six Engines

Alejo, K.; Fisher, S.; Kalluri, T.; More, B.; Rajgure, H.; Panda, P. K.; Korban, C.; Chung, C.

2026-08-11 biochemistry 10.64898/2026.08.11.744016 medRxiv
Top 0.1%
60.5%
Show abstract

Molecular docking and co-folding engines are widely used to prioritize compounds for wet-lab validation, yet their accuracy is known to vary substantially across protein targets for reasons that remain only qualitatively understood. Here we benchmark six docking and co-folding engines (RevDock, DiffDock, Boltz2, AutoDock-GPU, rDock, and PandaDock) across 14 protein families, evaluating scoring power, ranking power, docking power, and physical validity. Rather than treating engine performance as protein-family-specific, we classify all 14 families into six mechanistic groups according to which of four scoring-function simplifications, rigid receptor, pairwise additivity, fixed point charges, and implicit solvent, is most severely stressed by that familys binding site. This framework helps explain, rather than simply describe, where each engine succeeds or fails: RevDocks CNN rescoring layer mitigates the pairwise additivity and fixed-charge limitations relative to physics-only scoring, achieving the highest overall pose accuracy (73.3% of poses [≤] 2.0 [A] RMSD), while Boltz2s sequence-based co-folding bypasses the rigid-receptor assumption and achieves comparable affinity correlation (mean Pearson r {approx} 0.60 for both engines). PandaDock, run with expanded conformational sampling, matches RevDock on pose accuracy (72.1% of poses [≤] 2.0 [A], lowest median RMSD at 0.96 [A]) and exceeds AutoDock-GPU on affinity correlation (mean r = 0.460), indicating that the performance of a physics-based scoring function is limited as much by search adequacy as by the scoring function itself. These results suggest that engine selection for a docking or co-folding campaign should be guided less by an engines aggregate benchmark ranking and more by which of these four structural and physical characteristics dominate the target of interest.

3
PAM-DB: Revealing Protein Activation Mechanisms for Next-Generation Rational Drug Discovery

Zhu, X.; Li, X.; Hou, Y.; Zhou, R.; Yan, Y.; Warshel, A.; Bai, C.

2026-08-22 biophysics 10.64898/2026.08.20.745895 medRxiv
Top 0.1%
60.2%
Show abstract

Current rational drug design relies predominantly on computational (CADD/AIDD) methods that model binding thermodynamics and static conformations of target proteins, primarily in their inactive states. However, the kinetic parameters that govern experimental efficacy-such as catalytic turnover and signaling potency-are determined by molecular interactions with transition states (TS), intermediate states (IS), and the entire continuum of conformations along the least free-energy activation pathway. The absence of this dynamic dimension has fundamentally limited the predictive power and success rate of conventional structure-based approaches. Here, we present a structural database that systematically maps the complete activation trajectories of pharmaceutically relevant targets, encompassing TS, IS, and all connecting conformational ensembles. This resource offers multiple strategic advantages for drug discovery: enabling rational targeting of previously "undruggable" proteins, facilitating biased agonism/antagonism design, revealing cryptic allosteric sites in inactive conformations, identifying novel transient pockets along the activation route, rationalizing the mechanisms of existing drugs, predicting mutational effects on activation barriers, and prospectively forecasting drug resistance and off-target liabilities. We demonstrate the utility of this database through representative case studies and provide implementation guidelines for integration into existing discovery pipelines. More detailed information can be found at our website: https://www.momedpamdb.com/en.

4
MolJam: A Multidimensional Framework for Assessing Molecular Dataset Quality and Its Impact on Machine Learning

Wang, P.; Shi, Z.; Gao, X.; Zhou, R.

2026-08-25 bioinformatics 10.64898/2026.08.21.746384 medRxiv
Top 0.1%
60.0%
Show abstract

High-quality molecular datasets are essential for reliable machine learning in cheminformatics and bioinformatics, yet dataset quality is rarely assessed systematically and its relationship with downstream model performance remains poorly understood. Here, we present MolJam, an open-source framework for quantitative assessment of molecular dataset quality across five dimensions-structural integrity, data quality, experimental information quality, chemical space coverage, and data distribution-using 12 standardized metrics. Application of MolJam to 11 MoleculeNet and eight ChEMBL-derived datasets revealed widespread and heterogeneous quality issues, including undefined stereochemistry in up to 70.72% of molecules, inconsistent molecular representations, and contradictory labels. We next asked whether improving these quality metrics necessarily improves machine learning performance. Refinement of the ESOL and Lipophilicity datasets increased their MolJam quality scores but produced mixed effects on predictive performance, suggesting a competing influence of reduced dataset size. Controlled ablation experiments further demonstrated that both dataset quality and data quantity contribute to model performance and, notably, that retaining molecules with incomplete stereochemical information can outperform their removal when the resulting gain in data quantity offsets the quality penalty. Thus, molecular dataset curation cannot be reduced to maximizing data cleanliness alone but requires balancing multiple dimensions of data quality against information loss. MolJam provides a standardized framework for diagnosing molecular dataset limitations, comparing benchmark quality, and quantitatively evaluating how data curation decisions influence downstream machine learning.

5
Can SMILES be fragmented into a concatenable ordered sequence of retrosynthetically interesting string block ?

Reboul, E.; Prabakaran, H.; Baaden, M.; Waldispuhl, J.; Taly, A.

2026-08-26 bioinformatics 10.64898/2026.08.25.747180 medRxiv
Top 0.1%
57.8%
Show abstract

Molecules generated by deep learning models are often difficult to synthesize. Their synthetic accessibility can be improved with automated retrosynthetic analysis, which allows for identifying synthons. However, synthons in a SMILES can be scattered throughout the string depending on the path taken through the molecular graph used to generate the SMILES. We tested whether the ensemble of possible SMILES for a molecule can be used to generate a concatenable ordered sequence of string fragments (blocks) from SMILES that match potential synthons obtained through automated retrosynthetic analysis. We found that exhaustively sampling the SMILES space of a molecule improves the coverage of retrosynthetic breaks. We achieved full coverage of retrosynthetic bonds in string form for 85\% of the 1.9 million molecules in the MOSES dataset. Doing so allowed us to test our block SMILES in an unconditional de novo drug design test case with MolGPT and Monte Carlo Tree Search (MCTS). We found that using blocks as an LLM's token did degrade MolGPT performance due to the curse of dimensionality. However, using the SMILES selected by our blocking algorithm with the default SMILES tokenizer improved the reproduction of physico-chemical properties of samples and also improved uniqueness, novelty, and validity. The MCTS outperforms our MolGPT models in terms of validity and novelty. However, samples generated by the MCTS had physico-chemical properties that were further away from the MOSES baseline than the samples produced by molGPT, with an improved distribution of quantitative estimation of drug-likeness (QED).

6
Systematic Benchmarking of AI-Based Molecular Generation Models for Structure-Based Drug Design

Kumar, H.; Yang, Z.; Yu, Y.; Wen, J.; Kim, P.; Zhou, X.

2026-08-20 bioinformatics 10.64898/2026.08.14.744939 medRxiv
Top 0.1%
55.4%
Show abstract

Generative artificial intelligence is accelerating molecular design, yet the relative suitability of available models for different targets and stages of preclinical drug discovery remains unclear. Here we benchmarked 12 molecular generation and optimization methods across 176 curated protein-ligand systems spanning diverse therapeutic target classes, with experimentally validated ligands providing reference chemical space. The evaluated methods encompassed pocket-conditioned 3D generation, diffusion and flow-based modeling, autoregressive construction, reference-conditioned optimization and synthesis-aware design. Performance was assessed using operational robustness, chemical validity, uniqueness, molecular and scaffold diversity, quantitative estimate of drug-likeness, synthetic accessibility, docking, physicochemical and ADMET properties, and computational resource requirements. The results revealed architecture-dependent trade off such as receptor-conditioned methods exploited binding-pocket geometry, flow-based approaches enabled efficient sampling, reference-conditioned methods favored analogue generation, and synthesis-aware approaches improved chemical feasibility, but no method consistently optimized all criteria. To address the functional potential of generated molecules, we further developed a state-aware functional classifier (SAFC) that integrates molecular dynamics derived receptor ensembles, ensemble docking and protein ligand interaction graphs. SAFC provided dynamics-aware functional activity rankings for generated molecules that were partly complementary to docking, drug-likeness and synthetic accessibility scores. These findings support hybrid, stage specific deployment of generative models rather than reliance on any single architecture or evaluation metric. This study provides practical guidelines for generative AI based preclinical drug development processes.

7
A multi-agent molecular optimization framework leads to a rapid-recovery intravenous anesthetic candidate with an improved safety margin

Xue, Z.; Liu, X.

2026-08-20 bioinformatics 10.64898/2026.08.17.745149 medRxiv
Top 0.1%
54.0%
Show abstract

Lead optimization, the systematic refinement of therapeutic compounds through iterative structural modification, faces a dual challenge in modern drug discovery: navigating astronomically vast molecular design spaces while balancing conflicting demands on potency, pharmacokinetics, and safety. We present MASCOT (Multi-Agent SearCh for molecular OpTimization), a role-specialized multi-agent framework for molecular optimization. Integrated with a chemically constrained graph-editing search, MASCOT coordinates three specialized agents: a trade-off agent that reprioritizes competing objectives, a strategy agent that adapts how molecular edits are proposed, and a reflection agent that distills lessons from previous decisions. Computational experiments showed that MASCOT achieved the best performance over competing methods on six benchmark settings. On the SARS-CoV-2 main protease task, its mean docking-score improvement was 3.6 times that of the strongest baseline. Applied to the clinically used anesthetic remimazolam (RM), MASCOT prioritized RM-1, which showed a shorter liver microsomal half-life, higher brain exposure, and a larger therapeutic index than RM. Subsequent derivative design yielded RM-7. Extensive animal studies established RM-7 as a rapid-recovery intravenous anesthetic candidate with greater potency, faster functional recovery, a wider safety margin, and preserved flumazenil reversibility. These results demonstrate that multi-agent coordination can link adaptive molecular search to medicinal chemistry and experimental pharmacology.

8
A multimodal representation learning platform for accurate molecular ADMET prediction

Luo, Z.; Huang, D.; Shao, Y.; Yu, Q.; Li, Y.

2026-08-25 bioinformatics 10.64898/2026.08.24.746660 medRxiv
Top 0.1%
44.8%
Show abstract

Accurate ADMET prediction is essential for prioritizing compounds before costly experimental validation, yet ADMET tasks are highly heterogeneous. Properties such as solubility, permeability, protein binding, clearance, transporter activity and toxicity are governed by different molecular signals, ranging from local functional groups and physicochemical descriptors to bonded topology and three-dimensional geometry. Consequently, a single molecular representation or backbone is unlikely to be optimal across all ADMET tasks. We present Trimole-Hybrid, a task-wise multimodal framework that addresses ADMET heterogeneity by selecting or combining predictors built from complementary molecular representations. Trimole-Hybrid constructs a candidate pool of SMILES-, graph-, geometry-sensitive EPT/3D- and chemical descriptor-based predictors. For each task, Trimole-Hybrid selects the best-performing predictor to obtain the final prediction. On 22 Therapeutics Data Commons ADMET benchmarks, Trimole-Hybrid exceeded the public TDC top-1 methods on 10 tasks and ranked within the top 10 for 21 tasks. Ablation studies confirmed the contribution of both complementary multimodal molecular representations and task-specific ensemble strategies. In two small-molecule case studies, Trimole-Hybrid shows sensitivity to changes in essential functional motifs, suggesting its ability to capture ADMET-relevant molecular substructures.

9
Toward Robust Characterization of Dynamic Binding Pockets: Lessons from the HBV Capsid Assembly Modulator Site

Perez-Segura, C.; Scott, L. W.; Zlotnick, A.; Hadden-Perilla, J. A.

2026-08-10 biophysics 10.64898/2026.08.06.743403 medRxiv
Top 0.1%
40.6%
Show abstract

Protein function often depends on ligand binding pockets that fluctuate among conformational states, altering their size, shape, topology, and accessibility, yet quantitative comparison of these dynamic cavities remains challenging because their boundaries are often inherently ambiguous. The measure volinterior algorithm uses fuzzy-boundary detection to characterize enclosed molecular spaces; here, the hepatitis B virus (HBV) capsid assembly modulator (CAM) binding site is used as a model system to develop and validate a practical workflow for applying the method to dynamic protein binding pockets. The resulting methodology provides practical guidance for parameter selection and evaluation, establishes a standardized protocol for quantitative characterization of the HBV CAM pocket, and demonstrates robust, reproducible performance across conformational ensembles derived from molecular dynamics (MD) simulations. More broadly, this work provides a reproducible strategy for adapting measure volinterior to other dynamic binding pockets, enabling consistent comparison of pocket geometry among independent structural studies. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=105 SRC="FIGDIR/small/743403v1_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@d460b5org.highwire.dtl.DTLVardef@11915c2org.highwire.dtl.DTLVardef@1e37524org.highwire.dtl.DTLVardef@1fc1b7_HPS_FORMAT_FIGEXP M_FIG C_FIG

10
FlexAutoDock: A Flexible Platform for Automated Molecular Docking and Virtual Screening of Natural and Synthetic Compounds

Ahmed, M. F.; Faysal, M. F.; Sawad, K. M.; -E- Elahi, M. A.; Noor, T.; Kibria, M. K.; Hasan, M. M.; Mollah, M. N. H.

2026-08-19 bioinformatics 10.64898/2026.08.11.744098 medRxiv
Top 0.1%
31.4%
Show abstract

Drug discovery (DD) is a complex, time-consuming, and resource-intensive process that involves the identification of therapeutic targets, selection of bioactive compounds, and extensive experimental validation. The discovery of promising therapeutic compounds from large libraries of phytochemicals and synthetic molecules remains a major challenge in modern drug development. Screening millions of compounds through conventional experimental approaches requires substantial time, cost, and computational resources. In recent years, in- silico molecular docking has emerged as an important computational approach for predicting interactions between small molecules and target proteins, thereby helping researchers prioritize promising compounds for further investigation. Several molecular docking webservers, including iScreen, SwissDock, CB-Dock2, DockThor, and MTiOpenScreen, have been developed to support virtual screening studies. However, many currently available platforms still face some important limitations. Most existing tools lack integrated repositories of medicinal plant-derived phytochemicals and organism-derived bioactive compounds, automated mapping between plants and their associated phytochemicals, and flexible ligand retrieval using chemical names, SMILES strings, PubChem CIDs, or drug names. In addition, many platforms require extensive manual protein and ligand preparation, provide limited support for AlphaFold-predicted protein structures, and lack efficient large-scale multi-target virtual screening. Most existing docking platforms offer limited support for interactive inspection of docked protein-ligand complexes, often requiring users to download the results and analyse them using external molecular visualization software. To address these limitations, we developed FlexAutoDock, an automated cloud-based molecular docking platform that provides a unified environment for protein-ligand docking and large-scale virtual screening. Unlike existing web servers, FlexAutoDock integrates curated repositories of medicinal plant- derived phytochemicals, organism-derived bioactive compounds, and synthetic compounds from the ZINC database while supporting flexible ligand acquisition through medicinal plant or organism selection, chemical names, SMILES strings, PubChem CIDs, and drug-name queries. The platform further streamlines the docking workflow through automated protein structure retrieval from the Protein Data Bank and AlphaFold databases, receptor and ligand preparation, chain-specific protein selection, blind and site-specific docking, interactive visualization of predicted protein-ligand complexes, and scalable multi-target virtual screening. The resulting platform enables rapid, flexible, and large-scale virtual screening while simplifying the molecular docking workflow, providing researchers with an accessible computational resource for accelerating early-stage drug discovery. FlexAutoDock offers a fast, reliable, and accessible computational platform for molecular docking and virtual screening, freely available to the scientific community at http://103.99.177.82:3000/.

11
Scalable Extraction of Information on Protein-Protein Interactions using Topological Data Analysis

Mukherjee, A.; Park, B.; Malmstrom, A.; Cisewski-Kehe, J.; Van Lehn, R. C.; Zavala, V. M.

2026-08-09 bioinformatics 10.64898/2026.08.06.743405 medRxiv
Top 0.1%
31.2%
Show abstract

Protein-protein interactions (PPIs) govern a wide range of cellular functions. The ability to predict PPI interfaces from protein molecular surfaces is important for understanding protein function and enabling therapeutic discovery. While recent advances in structure-based learning, particularly molecular-surface geometric deep learning frameworks, have demonstrated that protein surfaces encode rich geometric and physicochemical information, such approaches often remain computationally intensive and data-hungry. Alternatively, topological data analysis (TDA) has emerged as a mathematically rigorous framework for extracting robust, multiscale shape information from complex data. In this work, we introduce a scalable TDA framework for extracting information on PPIs directly from localized protein surface patches. Our approach leverages multiscale topological descriptors, evaluated from patch-wise point cloud representations of protein mesh surfaces, combined with supervised machine learning models for interface prediction. On a full dataset of 3,362 proteins, the proposed approach substantially reduced computational cost relative to an established geometric deep learning method, MaSIF-site, decreasing preprocessing time from approximately 27 s/protein to 5-8 s/protein and total training time from approximately 6 h to 1-1.3 h. Importantly, this computational reduction is achieved while maintaining mean test area under the receiver operating characteristic curve (AUC) values of 0.76 and 0.77 for patch radii of 9 [A] and 12 [A], respectively, thus approaching the MaSIF-site test AUC of 0.84. Our results suggest that topology offers a scalable and computationally efficient approach for high-throughput extraction of information from complex biomolecular interfaces.

12
Evaluating Lightweight and Full Fine-Tuning Strategies Against Classical Machine Learning for Protein Function Prediction

Ab Ghani, N. S.; Matsushita, T.; Noguchi, T.; Kurumida, Y.; Kawada, S.; Ito, T.; Umetsu, M.; Saito, Y.

2026-08-07 bioinformatics 10.64898/2026.08.02.737389 medRxiv
Top 0.1%
31.0%
Show abstract

Motivation Protein language models (PLMs) have emerged as powerful tools for sequence-based prediction of protein function, yet systematic benchmarks comparing frozen embeddings, fine-tuning strategies like Low-Rank Adaptation (LoRA) and classical machine learning (ML) remain limited. We benchmarked four ML strategies: ML using amino acid descriptors (SL-AAFeat), ML using frozen embeddings from 20 PLMs across various pooling strategies (SL-Embed), full model fine-tuning (FT-Full) and LoRA-based fine-tuning (FT-LoRA). Performance was evaluated on the in-house VHH phage display dataset (VHH) for binding affinity prediction and the TAPE fluorescence dataset (FLS and FLS10) for mutational effect prediction. Results Model performance depended strongly on the dataset and adaptation strategy. Max pooling consistently improved embedding-based models, while amino acid descriptors remained competitive under specific datasets and resource constraints. Fine-tuning generally provided the highest predictive performance, but the advantage is not universal. Hyperparameter optimization significantly enhanced FT-LoRA, enabling it to outperform FT-Full on the VHH dataset with less than 10% model parameter adaptation. In contrast, FT-Full achieved the best performance on FLS and FLS10. Several medium-sized PLMs performed comparably to larger models, highlighting favorable performance-efficiency trade-offs. Overall, this paper presents a thorough review of PLM utilization strategies and practical recommendations for selecting suitable strategies based on dataset characteristics and available computational resources. Availability The source code used in this manuscript is available in a Zenodo repository at https://doi.org/10.5281/zenodo.21466255.

13
The first OpenBind release: An open experimental structure-affinity dataset and benchmark for structure-based AI

Nelen, J.; Khan, O.; Adams, E.; Aschenbrenner, J. C.; Thompson, W.; Ebrahim, A.; Capkin, E.; Vallee, C.; OpenBind, ; Shotton, E. J.; Griffen, E. J.; Chodera, J. D.; Deane, C. M.; von Delft, F.; AlQuraishi, M.; Imrie, F.

2026-09-01 bioinformatics 10.64898/2026.08.27.747600 medRxiv
Top 0.2%
26.9%
Show abstract

High-quality experimental datasets that link protein-ligand structures with binding affinity data are essential for developing and evaluating structure-based machine learning methods. To help address this need, we established OpenBind as an open-science initiative to generate large-scale experimental datasets for structure-based AI and molecular discovery. Here, we describe the first public OpenBind release, which, to the best of our knowledge, is the largest public single-target experimental structure-affinity dataset. The dataset focuses on enteroviral 2A protease, comprising 925 crystallographic binding events from 699 compounds and associated affinity measurements for 601 compounds. It combines structures from an initial fragment screen and follow-on molecules, together with affinity data, linking experimentally determined protein-ligand binding modes to biophysical measurements within a coherent antiviral discovery campaign. We used this dataset to evaluate protein-ligand structure prediction, binding-affinity prediction, and virtual screening using representative structure-based methods, including docking and cofolding. This exposed several challenges that are central to practical structure-based modelling: docking performance depends strongly on binding-pocket conformation, poses are difficult to rank, and structure-based affinity prediction remains challenging. Fine-tuning OpenFold3-p2 on the fragment-screen structures substantially improved pose prediction and virtual screening for related follow-on compounds, demonstrating how early-stage experimental structures can support target-specific model adaptation.

14
Benchmarking AI-generated structural ensembles of membrane proteins against physics-based modelling

Clifton, B. R.; Grieve, A. G.; Corey, R. A.

2026-08-09 biophysics 10.64898/2026.08.08.743655 medRxiv
Top 0.2%
26.9%
Show abstract

Proteins dynamically switch between a continuum of interconverting conformational states, and understanding these structural dynamics is important for understanding protein function and for developing therapeutics. Molecular dynamics (MD) simulations can provide insight into protein conformational ensembles, but sampling rare conformational states can require substantial computational resources. The recent development of AI-based approaches for generating protein conformational ensembles, such as the Biomolecular Emulator (BioEmu), offers a potential alternative, although it remains unclear whether these approaches can accurately capture the conformational landscapes, especially for special cases such as membrane proteins. Here, we assess the ability of BioEmu to model the conformational dynamics of a model membrane protein, the bacterial rhomboid intramembrane proteases GlpG. We find that BioEmu generates a range of conformations corresponding to both open and closed states of the rhomboid lateral gate, including states associated with different stages of the catalytic cycle. These conformations broadly correspond to states sampled during microsecond-timescale MD simulations, although BioEmu does not reproduce the full conformational landscape observed using MD. BioEmu also samples substantial conformational heterogeneity within the soluble domains of rhomboids, which are highly flexible and poorly represented in experimental structures. Overall, our findings demonstrate that BioEmu can generate plausible conformational ensembles for relatively large, six-and seven-pass membrane proteins, sampling rare states at a fraction of the computational cost of conventional MD simulations. These results suggest that AI-based ensemble generation could provide an accessible approach for exploring membrane protein dynamics and complement conventional molecular modelling approaches.

15
PandaDock: An Open-Source Molecular Docking Platform with Flexible-Ligand Search and Equivariant Neural Scoring

Panda, P. K.

2026-08-20 bioinformatics 10.64898/2026.08.19.745667 medRxiv
Top 0.2%
26.9%
Show abstract

We present PandaDock, an open-source molecular docking platform implementing flexible-ligand conformational search with analytic gradients, a precomputed affinity grid engine, specialized modules for induced-fit, metal-coordination and tethered docking, and an SE(3)-equivariant graph neural network scoring function trained at scale. Ligand flexibility is represented as a torsion tree and pose parameters are optimized by Monte Carlo with Metropolis acceptance refined by L-BFGS, with rotational gradients obtained in closed form through the derivative of the SO(3) exponential map rather than by finite differences. Affinity grids are built by a blocked neighbor-selection scheme that is exact and 5.6-9.7x faster than dense evaluation, and may be cached across ligands sharing a receptor and site, reducing a six-ligand series from 29.3 s to 10.4 s. On 814 protein-ligand complexes spanning 14 target families, PandaDock recovers a pose within 2 Angstroms of the crystal geometry in 33.7% of cases at rank 1 and in 57.0% of cases within the returned ensemble. The GNN scoring function is trained on 741,706 co-folded complexes from SAIR under target-disjoint splits, reaching a Pearson r of 0.407 on 90,219 held-out complexes and transferring to 202 independent crystal structures with measured Ki, Kd, IC50 or EC50 at r = 0.467. We report the model against three controls, a target-mean predictor, a ligand-descriptor-only baseline, and within-target correlations, and document both where it performs and where it does not, including its unsuitability for pose rescoring. On an independent 30-compound series against a single GABAA receptor target, PandaDock's empirical scoring function ranks 8th of 25 methods evaluated, ahead of every AutoDock Vina and Vinardo configuration tested, while the GNN scores below Vina, consistent with the within-target ceiling identified on SAIR. At full scale on the PDBbind v2020 refined set (n = 4,640, native crystal poses), the fully independent SAIR model reaches r = 0.531, and a dedicated model trained on PDBbind alone under a target-disjoint split reaches r = 0.690 on its own held-out test complexes, the strongest evidence in this work that PandaDock's affinity predictions generalize. PandaDock is distributed under an open-source license at https://github.com/pritampanda15/PandaDock with a complete command-line interface and a reproducible benchmarking harness.

16
Mechanistic Dissection of Entropic Penalty upon Ligand Binding and Molecular Flexibility via Molecular Dynamics Simulations and Machine Learning

Hung, T. I.; Vig, E.; Chang, C.-e.

2026-08-20 biophysics 10.64898/2026.08.18.745526 medRxiv
Top 0.2%
26.4%
Show abstract

Molecular flexibility governs how molecules behave, reorganize, and respond to their environment. Although experiments measure molar entropy for small molecules and molecular dynamics (MD) simulations capture molecular motions, quantifying configuration entropy and the concerted internal motions such as torsion rotations, angle bending, and their couplings are central to understanding thermodynamic behavior but remains challenging. To dissect these contributions, we used MD trajectories and developed an internal coordinate PC-entropy (iPC-entropy) method to probe the origins of entropy and reveal how specific motions shape the thermodynamic landscape. The studies accurately captured molar entropy, identified key torsional motions as major contributors, and uncovered a critical angle-torsion coupling in which angle bending was strongly correlated with torsional rotation, a coupling that increases nonlinearly with molecular size. Evaluating entropic changes upon protein-ligand binding reveals that dominant entropic penalty arises from ligand dihedral rigidification rather than protein reorganization and highlights the specific dihedral rotations that become restricted. We also suggest systematic corrections for approaches considering solely rotamers to reliably reproduce the relative entropic penalty in computer-aided drug discovery. Together, our findings elucidate the molecular origins of entropy and entropy changes. In addition, we can quantify and illustrate the internal motions that strongly shape binding thermodynamics, thereby offering mechanistic insights to guide drug development.

17
Ab initio side-chain sampling with PUD+ enables high-fidelity protein dynamics across AI-driven and classical simulations

Wu, D.; Wang, T.

2026-08-11 biophysics 10.64898/2026.08.10.743906 medRxiv
Top 0.2%
26.4%
Show abstract

The fidelity of molecular dynamics (MD) simulations fundamentally depends on the quality and coverage of the ab initio data used to parameterize the underlying force field, yet the role of side-chain conformational space remains insufficiently explored. In this study, we systematically investigate how comprehensive ab initio sampling of dipeptide conformations--specifically targeting side-chain degrees of freedom--impacts force field accuracy and MD simulation predictive power. We present the Protein Unit Dataset Plus (PUD+), a 40-million-conformation quantum mechanical dataset featuring unprecedented coverage of both backbone and side-chain conformational space. Machine learning force fields trained on PUD+ and integrated into AI2BMD simulations demonstrate superior energy and force prediction accuracy, capturing high-fidelity protein folding dynamics and the conformational flexibility of long-side-chain systems. Furthermore, leveraging PUD+ to reparameterize the CMAP term of the classical ff19SB force field markedly improves the description of intrinsically disordered protein (IDP) dynamics and IDP-ligand binding. Collectively, these results demonstrate that ab initio sampling of dipeptide side-chain conformations enables high-fidelity modeling of protein dynamics across both AI-driven and classical simulation paradigms.

18
A transition state-like acylenzyme conformation distinguishes carbapenemase activity in class A β-lactamases

Beer, M.; Spencer, J.; Mulholland, A. J.

2026-09-01 biochemistry 10.64898/2026.08.31.748333 medRxiv
Top 0.2%
26.0%
Show abstract

Carbapenems are the most potent {beta}-lactams, key antibiotics for healthcare-associated infections by Gram-negative bacteria and evade hydrolysis by most {beta}-lactamases, but are increasingly threatened by emergence of enzymes exhibiting hydrolytic activity towards them. Of the four recognised {beta}-lactamase subclasses, class A (active-site serine enzymes that hydrolyse {beta}-lactams via a covalent acylenzyme intermediate) is the most widely disseminated and, while the majority of such enzymes react with carbapenems to form long-lasting acylenzyme complexes, several possess carbapenem-hydrolyzing activity (carbapenemases). Here, we investigate the basis for these differences in a panel of class A {beta}-lactamases using molecular dynamics (MD) simulations of the respective acylenzyme complexes and tetrahedral intermediates (TI). The simulations reveal multiple features associated with catalytic activity across the spectrum of enzymes tested, including more extensive interactions of the carbapenem acylenzyme carbonyl and generally increased lifetimes of active site water molecules positioned for deacylation. Analysis of the dynamic trajectories shows carbapenemases to have reduced root mean-squared fluctuation (RMSF) differences between the acylenzyme and TI, that are not limited to the active site, indicating that the acylenzyme complex is pre-organised for reaction in carbapenemases but not in carbapenem-inhibited enzymes. Similarly, Principal Component Analysis (PCA) of acylenzyme and TI dynamics shows greater overlap between the two states in carbapenemases, providing further evidence for acylenzyme pre-organisation. Such simulations may represent an effective computational assay able to identify enzymes with carbapenemase activity at relatively modest computational cost.

19
Gaussian Accelerated Molecular Dynamics in GROMACS

Yang, Y.

2026-08-10 biochemistry 10.64898/2026.08.10.743837 medRxiv
Top 0.2%
22.7%
Show abstract

Gaussian accelerated molecular dynamics (GaMD) enhances conformational sampling by adding a smooth boost potential without requiring predefined collective variables, but an engine-integrated implementation has not been available in GROMACS. Here, we implement total-, dihedral-, and dual-boost GaMD in GROMACS 2025.4, including staged energy-statistics collection, GPU-based bias evaluation and force scaling, restart support, and outputs required for cumulant-based free-energy reweighting. The implementation was evaluated using four benchmark systems spanning conformational free energies, protein folding, and ligand recognition. For alanine dipeptide, a reweighted 100 ns GaMD trajectory recovered the major free-energy basins and rotational barriers in overall agreement with a 1000 ns conventional MD simulation. For chignolin and TC5b, all three independent trajectories for each system sampled native-like folded states from extended conformations within 300 ns and 1 s, respectively; the best TC5b structure had a minimum backbone RMSD of 0.03 nm from the experimental structure. In the benzene-T4 lysozyme system, two of five independent 500 ns trajectories captured both ligand binding and dissociation, yielding a bound pose with a minimum ligand RMSD of 0.06 nm from the crystal structure. Across all four systems, the boost-potential distributions were approximately Gaussian, and second-order cumulant reweighting resolved the expected conformational and binding free-energy basins. These results demonstrate that GROMACS-GaMD provides a practical, GPU-enabled, collective-variable-free enhanced-sampling framework for biomolecular free-energy calculations, protein folding, and ligand-binding studies.

20
Application of 3D Zernike Descriptors in Antibody Structural Clustering and Repurposing

de Almeida, D. d. S.; Albuquerque, A. O.; Peixoto Lima, A. M.; Gaieta, E. M.; Souza, J. S.; dos Santos-Costa, A. H.; de Andrade, L. M.; Sampaio, J. V.; Sartori, G. R.; Silva, e. J. H. M. d.

2026-08-19 bioinformatics 10.64898/2026.08.12.744489 medRxiv
Top 0.2%
22.6%
Show abstract

Antibodies generally exhibit high specificity for their cognate epitopes, but structural and physicochemical similarities between distinct epitopes can enable an antibody to recognize different antigens, resulting in cross-reactivity. This property can be exploited for antibody repurposing. To identify epitopes that share such similarities, both sequence- and structure-based approaches can be employed. In this context, 3D Zernike descriptors provide a compact representation of protein surface geometry as numerical feature vectors, enabling quantitative comparisons independently of structural alignment and orientation. Thus, this study aimed to evaluate the application of 3D Zernike descriptors for the structural clustering of antibodies and epitopes and to explore their use in antibody repurposing for the recognition of new targets. To this end, antibody binding sites previously associated with recognition of similar epitopes were analyzed at different structural levels, considering the CDRs, CDRH3, and complete paratopes. Surface similarity was subsequently quantified by calculating the Euclidean distance between their corresponding 3D Zernike feature vectors. Performance was benchmarked against SPACE2. Additionally, different distance thresholds were evaluated based on their ability to recover antibody pairs recognizing the same epitope. The paratope-based approach provided the best balance between the number of identified pairs and precision at a distance threshold of 2.7, whereas epitope clustering showed robust performance up to a distance of 3.0. At these thresholds, the 3D Zernike descriptors identified a greater number of functional pairs than SPACE2 while maintaining comparable precision and identifying complementary sets of antibody pairs.. BTaken together, these findings support the use of 3D Zernike descriptors for structural clustering of antibodies and epitopes and for guiding antibody repurposing G, a highly lethal zoonotic pathogen. Structural screening identified three antibodies with epitopes similar to the NiV target that also showed a consistent binding preference for the target epitope in molecular docking assays. Notably, one candidate, originally directed against a SARS-CoV-2 epitope, formed a stable complex with the NiV epitope, remaining within the 5 [A] RMSD threshold during heated molecular dynamics simulations and emerging as a potential cross-reactive candidate.These results support the use of this computational framework for biopharmaceutical discovery against emerging targets. Taken together, these findings support the use of 3D Zernike descriptors for structural clustering of antibodies and epitopes and for guiding antibody repurposing.